Skip to content

[ExecuTorch][Vulkan] Tests for et_vk.q4gsw_requant#20946

Merged
meta-codesync[bot] merged 12 commits into
gh/JCNTH/85/basefrom
gh/JCNTH/85/head
Jul 21, 2026
Merged

[ExecuTorch][Vulkan] Tests for et_vk.q4gsw_requant#20946
meta-codesync[bot] merged 12 commits into
gh/JCNTH/85/basefrom
gh/JCNTH/85/head

Conversation

@JCNTH

@JCNTH JCNTH commented Jul 14, 2026

Copy link
Copy Markdown
Contributor

Stack from ghstack (oldest at bottom):

Correctness tests for the Vulkan et_vk.q4gsw_requant kernel (stacked above the op diff).

Coverage: the golden codes are computed with ATen (round/clamp, zero-scale -> code 8), mirroring quant_nibble, then packed into the expected W_4X8 int buffer with a small bit-packing reference (data-reshaping only, no hand-rolled math). The kernel output is compared int-for-int against that buffer, which locks the exact byte layout the forward reads. The latent is built as code * scale so round() is unambiguous (no .5 tie-break divergence).

Cases:

  • test_tile_aligned — single group, tile-aligned N/K.
  • test_grouped — multiple quantization groups along K.
  • test_odd_n4N % 8 != 0 (odd N4 -> padded stride + bias-zero OOB tile).
  • test_zero_scale — a zero scale must yield code 8, not a divide-by-zero.

Also wires q4gsw_requant_test into targets.bzl + CMakeLists.txt.
@exported-using-ghexport

Differential Revision: D111797526

Differential Revision: D111797526

[ghstack-poisoned]
@pytorch-bot pytorch-bot Bot added the module: vulkan Issues related to the Vulkan delegate and code under backends/vulkan/ label Jul 14, 2026
@pytorch-bot

pytorch-bot Bot commented Jul 14, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20946

Note: Links to docs will display an error until the docs builds have been completed.

❌ 2 New Failures, 1 Unrelated Failure

As of commit 5ee530f with merge base 21554e5 (image):

NEW FAILURES - The following jobs have failed:

FLAKY - The following job failed but was likely due to flakiness present on trunk:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

This was referenced Jul 14, 2026
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
@meta-codesync
meta-codesync Bot merged commit 86b8942 into gh/JCNTH/85/base Jul 21, 2026
186 of 189 checks passed
@meta-codesync
meta-codesync Bot deleted the gh/JCNTH/85/head branch July 21, 2026 16:25
@meta-codesync
meta-codesync Bot temporarily deployed to cherry-pick-bot July 21, 2026 16:26 Inactive
JCNTH added a commit that referenced this pull request Jul 21, 2026
Pull Request resolved: #20946

**Correctness tests for the Vulkan `et_vk.q4gsw_requant` kernel** (stacked above the op diff).

**Coverage:** the golden codes are computed with ATen (`round`/`clamp`, zero-scale -> code 8), mirroring `quant_nibble`, then packed into the expected W_4X8 int buffer with a small bit-packing reference (data-reshaping only, no hand-rolled math). The kernel output is compared int-for-int against that buffer, which locks the exact byte layout the forward reads. The latent is built as `code * scale` so `round()` is unambiguous (no `.5` tie-break divergence).

Cases:
- `test_tile_aligned` — single group, tile-aligned N/K.
- `test_grouped` — multiple quantization groups along K.
- `test_odd_n4` — `N % 8 != 0` (odd N4 -> padded stride + bias-zero OOB tile).
- `test_zero_scale` — a zero scale must yield code 8, not a divide-by-zero.

Also wires `q4gsw_requant_test` into `targets.bzl` + `CMakeLists.txt`.
ghstack-source-id: 405059122
@exported-using-ghexport

Differential Revision: [D111797526](https://our.internmc.facebook.com/intern/diff/D111797526/)
JCNTH added a commit that referenced this pull request Jul 21, 2026
Pull Request resolved: #20946

**Correctness tests for the Vulkan `et_vk.q4gsw_requant` kernel** (stacked above the op diff).

**Coverage:** the golden codes are computed with ATen (`round`/`clamp`, zero-scale -> code 8), mirroring `quant_nibble`, then packed into the expected W_4X8 int buffer with a small bit-packing reference (data-reshaping only, no hand-rolled math). The kernel output is compared int-for-int against that buffer, which locks the exact byte layout the forward reads. The latent is built as `code * scale` so `round()` is unambiguous (no `.5` tie-break divergence).

Cases:
- `test_tile_aligned` — single group, tile-aligned N/K.
- `test_grouped` — multiple quantization groups along K.
- `test_odd_n4` — `N % 8 != 0` (odd N4 -> padded stride + bias-zero OOB tile).
- `test_zero_scale` — a zero scale must yield code 8, not a divide-by-zero.

Also wires `q4gsw_requant_test` into `targets.bzl` + `CMakeLists.txt`.
ghstack-source-id: 405059122
@exported-using-ghexport

Differential Revision: [D111797526](https://our.internmc.facebook.com/intern/diff/D111797526/)
JCNTH added a commit that referenced this pull request Jul 21, 2026
Pull Request resolved: #20946

**Correctness tests for the Vulkan `et_vk.q4gsw_requant` kernel** (stacked above the op diff).

**Coverage:** the golden codes are computed with ATen (`round`/`clamp`, zero-scale -> code 8), mirroring `quant_nibble`, then packed into the expected W_4X8 int buffer with a small bit-packing reference (data-reshaping only, no hand-rolled math). The kernel output is compared int-for-int against that buffer, which locks the exact byte layout the forward reads. The latent is built as `code * scale` so `round()` is unambiguous (no `.5` tie-break divergence).

Cases:
- `test_tile_aligned` — single group, tile-aligned N/K.
- `test_grouped` — multiple quantization groups along K.
- `test_odd_n4` — `N % 8 != 0` (odd N4 -> padded stride + bias-zero OOB tile).
- `test_zero_scale` — a zero scale must yield code 8, not a divide-by-zero.

Also wires `q4gsw_requant_test` into `targets.bzl` + `CMakeLists.txt`.
ghstack-source-id: 405059122
@exported-using-ghexport

Differential Revision: [D111797526](https://our.internmc.facebook.com/intern/diff/D111797526/)
JCNTH added a commit that referenced this pull request Jul 21, 2026
Pull Request resolved: #20946

**Correctness tests for the Vulkan `et_vk.q4gsw_requant` kernel** (stacked above the op diff).

**Coverage:** the golden codes are computed with ATen (`round`/`clamp`, zero-scale -> code 8), mirroring `quant_nibble`, then packed into the expected W_4X8 int buffer with a small bit-packing reference (data-reshaping only, no hand-rolled math). The kernel output is compared int-for-int against that buffer, which locks the exact byte layout the forward reads. The latent is built as `code * scale` so `round()` is unambiguous (no `.5` tie-break divergence).

Cases:
- `test_tile_aligned` — single group, tile-aligned N/K.
- `test_grouped` — multiple quantization groups along K.
- `test_odd_n4` — `N % 8 != 0` (odd N4 -> padded stride + bias-zero OOB tile).
- `test_zero_scale` — a zero scale must yield code 8, not a divide-by-zero.

Also wires `q4gsw_requant_test` into `targets.bzl` + `CMakeLists.txt`.
ghstack-source-id: 405059122
@exported-using-ghexport

Differential Revision: [D111797526](https://our.internmc.facebook.com/intern/diff/D111797526/)
JCNTH added a commit that referenced this pull request Jul 21, 2026
Pull Request resolved: #20946

**Correctness tests for the Vulkan `et_vk.q4gsw_requant` kernel** (stacked above the op diff).

**Coverage:** the golden codes are computed with ATen (`round`/`clamp`, zero-scale -> code 8), mirroring `quant_nibble`, then packed into the expected W_4X8 int buffer with a small bit-packing reference (data-reshaping only, no hand-rolled math). The kernel output is compared int-for-int against that buffer, which locks the exact byte layout the forward reads. The latent is built as `code * scale` so `round()` is unambiguous (no `.5` tie-break divergence).

Cases:
- `test_tile_aligned` — single group, tile-aligned N/K.
- `test_grouped` — multiple quantization groups along K.
- `test_odd_n4` — `N % 8 != 0` (odd N4 -> padded stride + bias-zero OOB tile).
- `test_zero_scale` — a zero scale must yield code 8, not a divide-by-zero.

Also wires `q4gsw_requant_test` into `targets.bzl` + `CMakeLists.txt`.
ghstack-source-id: 405059122
@exported-using-ghexport

Differential Revision: [D111797526](https://our.internmc.facebook.com/intern/diff/D111797526/)
JCNTH added a commit that referenced this pull request Jul 21, 2026
Pull Request resolved: #20946

**Correctness tests for the Vulkan `et_vk.q4gsw_requant` kernel** (stacked above the op diff).

**Coverage:** the golden codes are computed with ATen (`round`/`clamp`, zero-scale -> code 8), mirroring `quant_nibble`, then packed into the expected W_4X8 int buffer with a small bit-packing reference (data-reshaping only, no hand-rolled math). The kernel output is compared int-for-int against that buffer, which locks the exact byte layout the forward reads. The latent is built as `code * scale` so `round()` is unambiguous (no `.5` tie-break divergence).

Cases:
- `test_tile_aligned` — single group, tile-aligned N/K.
- `test_grouped` — multiple quantization groups along K.
- `test_odd_n4` — `N % 8 != 0` (odd N4 -> padded stride + bias-zero OOB tile).
- `test_zero_scale` — a zero scale must yield code 8, not a divide-by-zero.

Also wires `q4gsw_requant_test` into `targets.bzl` + `CMakeLists.txt`.
ghstack-source-id: 405059122
@exported-using-ghexport

Differential Revision: [D111797526](https://our.internmc.facebook.com/intern/diff/D111797526/)
JCNTH added a commit that referenced this pull request Jul 21, 2026
Pull Request resolved: #20946

**Correctness tests for the Vulkan `et_vk.q4gsw_requant` kernel** (stacked above the op diff).

**Coverage:** the golden codes are computed with ATen (`round`/`clamp`, zero-scale -> code 8), mirroring `quant_nibble`, then packed into the expected W_4X8 int buffer with a small bit-packing reference (data-reshaping only, no hand-rolled math). The kernel output is compared int-for-int against that buffer, which locks the exact byte layout the forward reads. The latent is built as `code * scale` so `round()` is unambiguous (no `.5` tie-break divergence).

Cases:
- `test_tile_aligned` — single group, tile-aligned N/K.
- `test_grouped` — multiple quantization groups along K.
- `test_odd_n4` — `N % 8 != 0` (odd N4 -> padded stride + bias-zero OOB tile).
- `test_zero_scale` — a zero scale must yield code 8, not a divide-by-zero.

Also wires `q4gsw_requant_test` into `targets.bzl` + `CMakeLists.txt`.
ghstack-source-id: 405059122
@exported-using-ghexport

Differential Revision: [D111797526](https://our.internmc.facebook.com/intern/diff/D111797526/)
JCNTH added a commit that referenced this pull request Jul 21, 2026
Pull Request resolved: #20946

**Correctness tests for the Vulkan `et_vk.q4gsw_requant` kernel** (stacked above the op diff).

**Coverage:** the golden codes are computed with ATen (`round`/`clamp`, zero-scale -> code 8), mirroring `quant_nibble`, then packed into the expected W_4X8 int buffer with a small bit-packing reference (data-reshaping only, no hand-rolled math). The kernel output is compared int-for-int against that buffer, which locks the exact byte layout the forward reads. The latent is built as `code * scale` so `round()` is unambiguous (no `.5` tie-break divergence).

Cases:
- `test_tile_aligned` — single group, tile-aligned N/K.
- `test_grouped` — multiple quantization groups along K.
- `test_odd_n4` — `N % 8 != 0` (odd N4 -> padded stride + bias-zero OOB tile).
- `test_zero_scale` — a zero scale must yield code 8, not a divide-by-zero.

Also wires `q4gsw_requant_test` into `targets.bzl` + `CMakeLists.txt`.
ghstack-source-id: 405059122
@exported-using-ghexport

Differential Revision: [D111797526](https://our.internmc.facebook.com/intern/diff/D111797526/)
JCNTH added a commit that referenced this pull request Jul 21, 2026
Pull Request resolved: #20946

**Correctness tests for the Vulkan `et_vk.q4gsw_requant` kernel** (stacked above the op diff).

**Coverage:** the golden codes are computed with ATen (`round`/`clamp`, zero-scale -> code 8), mirroring `quant_nibble`, then packed into the expected W_4X8 int buffer with a small bit-packing reference (data-reshaping only, no hand-rolled math). The kernel output is compared int-for-int against that buffer, which locks the exact byte layout the forward reads. The latent is built as `code * scale` so `round()` is unambiguous (no `.5` tie-break divergence).

Cases:
- `test_tile_aligned` — single group, tile-aligned N/K.
- `test_grouped` — multiple quantization groups along K.
- `test_odd_n4` — `N % 8 != 0` (odd N4 -> padded stride + bias-zero OOB tile).
- `test_zero_scale` — a zero scale must yield code 8, not a divide-by-zero.

Also wires `q4gsw_requant_test` into `targets.bzl` + `CMakeLists.txt`.
ghstack-source-id: 405059122
@exported-using-ghexport

Differential Revision: [D111797526](https://our.internmc.facebook.com/intern/diff/D111797526/)
JCNTH added a commit that referenced this pull request Jul 21, 2026
Pull Request resolved: #20946

**Correctness tests for the Vulkan `et_vk.q4gsw_requant` kernel** (stacked above the op diff).

**Coverage:** the golden codes are computed with ATen (`round`/`clamp`, zero-scale -> code 8), mirroring `quant_nibble`, then packed into the expected W_4X8 int buffer with a small bit-packing reference (data-reshaping only, no hand-rolled math). The kernel output is compared int-for-int against that buffer, which locks the exact byte layout the forward reads. The latent is built as `code * scale` so `round()` is unambiguous (no `.5` tie-break divergence).

Cases:
- `test_tile_aligned` — single group, tile-aligned N/K.
- `test_grouped` — multiple quantization groups along K.
- `test_odd_n4` — `N % 8 != 0` (odd N4 -> padded stride + bias-zero OOB tile).
- `test_zero_scale` — a zero scale must yield code 8, not a divide-by-zero.

Also wires `q4gsw_requant_test` into `targets.bzl` + `CMakeLists.txt`.
ghstack-source-id: 405059122
@exported-using-ghexport

Differential Revision: [D111797526](https://our.internmc.facebook.com/intern/diff/D111797526/)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. meta-exported module: vulkan Issues related to the Vulkan delegate and code under backends/vulkan/

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants